WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse
State of the Art Pipeline Parallel Algorithm Research Paper
While I was working on Kimi K3, I decided to explore Pipeline Parallelism (PP) for improving prefill throughput. I had already implemented Kimi Linear Cache (KDA Cache) in TensorRT-LLM much before NVIDIA did, which reached a very high cache hit rate consistently in the industry.
I found a lot of bottlenecks in making this work for PP, just due to the hybrid cache mechanisms which needed multiple syncs in the code. I decided to completely redesign the PP code in TensorRT-LLM to work better with the hybrid cache.
I tried to envision the pipeline stages from a different perspective, that once data enters a stage, we shouldn't have to do anything for it after that. I.e., it should just move like continuous waves through the pipeline. This led me to extract most of the host-side and scheduler logic out of the main executor loop, and keep it outside. Not only did this make the algorithm faster, but it also made my work very flexible
At the end, I was able to even beat vLLM, SGLang on K3, and the existing TRT-LLM PP implementations by a significant margins. WavePP is up to 2x faster than TRT-LLM base, and 20% better than the second best PP implementation (even at concurrency 8) for K3.
There's a lot of interesting ideas I developed here to make prefix cache usage efficient in a pipeline with minimal synchronization, and asynchronous request admission. You can read about it here:
I hope you enjoyed reading the paper!!
Any feedback is greatly appreciated!